Goto

Collaborating Authors

 muse voice transcribe


Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time

Engadget

Meta has introduced its first real-time audio model, Muse Voice Transcribe. The model can handle dictation and transcription for more than 20 speakers and can seamlessly handle multiple languages at once, Meta says. Meta CEO Mark Zuckerberg, who recently returned to X after three years of not posting on the platform, shared an example of the model's ability to handle multiple speakers and languages at once. In the video, the transcription is able to automatically distinguish between multiple speakers and switch between languages. It's even able to pick up on "code-switching" and transcribe sentences that use words from multiple languages.